Listening while Speaking: Speech Chain by Deep Learning

机译：口语聆听：深度学习的语音链

代理获取

本网站仅为用户提供外文OA文献查询和代理获取服务，本网站没有原文。下单后我们将采用程序或人工为您竭诚获取高质量的原文，但由于OA文献来源多样且变更频繁，仍可能出现获取不到、文献不完整或与标题不符等情况，如果获取不到我们将提供退款服务。请知悉。

页面导航

摘要
著录项
相似文献
相关主题

摘要

Despite the close relationship between speech perception and production,research in automatic speech recognition (ASR) and text-to-speech synthesis(TTS) has progressed more or less independently without exerting much mutualinfluence on each other. In human communication, on the other hand, aclosed-loop speech chain mechanism with auditory feedback from the speaker'smouth to her ear is crucial. In this paper, we take a step further and developa closed-loop speech chain model based on deep learning. Thesequence-to-sequence model in close-loop architecture allows us to train ourmodel on the concatenation of both labeled and unlabeled data. While ASRtranscribes the unlabeled speech features, TTS attempts to reconstruct theoriginal speech waveform based on the text from ASR. In the opposite direction,ASR also attempts to reconstruct the original text transcription given thesynthesized speech. To the best of our knowledge, this is the first deeplearning model that integrates human speech perception and productionbehaviors. Our experimental results show that the proposed approachsignificantly improved the performance more than separate systems that wereonly trained with labeled data.

机译：尽管语音感知和产生之间有着密切的关系，但是自动语音识别（ASR）和文本到语音合成（TTS）的研究或多或少地独立进行，彼此之间没有很大的相互影响。另一方面，在人类交流中，具有从说话者的嘴到她的耳朵的听觉反馈的闭环语音链机制至关重要。在本文中，我们将进一步采取措施，并开发基于深度学习的闭环语音链模型。闭环体系结构中的按序序列模型使我们可以在标记数据和未标记数据的串联上训练模型。当ASR转录未标记的语音特征时，TTS尝试根据ASR的文本重建原始语音波形。在相反的方向上，ASR还尝试在给定合成语音的情况下重建原始文本转录。据我们所知，这是第一个将人类语音感知和生产行为整合在一起的深度学习模型。我们的实验结果表明，与仅使用标记数据进行训练的单独系统相比，所提出的方法显着提高了性能。

著录项

作者
Tjandra, Andros; Sakti, Sakriani; Nakamura, Satoshi;
展开▼
作者单位

展开▼
年度 2017
总页数
原文格式 PDF
正文语种
中图分类

相似文献

外文文献
中文文献
专利

1. SpeakerBeam: A New Deep Learning Technology for Extracting Speech of a Target Speaker Based on the Speaker’s Voice Characteristics [J] . Marc Delcroix, Katerina Zmolikova, Keisuke Kinoshita, NTT Technical Review . 2018,第11期

机译：SpeakerBeam：一种新的深度学习技术，用于根据说话者的语音特征提取目标说话者的语音
2. An Efficient Deep Learning Based Method for Speech Assessment of Mandarin-Speaking Aphasic Patients [J] . Mahmoud Seedahmed S., Kumar Akshay, Tang Yiting, Biomedical and Health Informatics, IEEE Journal of . 2020,第11期

机译：一种高效的基于深度学习的讲话评估方法，讲话者的失性患者
3. Developing AI that Pays Attention to Who You Want to Listen to: Deep-learning-based Selective Hearing with SpeakerBeam [J] . Marc Delcroix, Tsubasa Ochiai, Hiroshi Sato, NTT Technical Review . 2021,第9期

机译：开发一个人注意到要听取谁的人：基于深度学习的选择性听力与扬声器
4. Listening while speaking: Speech chain by deep learning [C] . Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 2017 IEEE Automatic Speech Recognition and Understanding Workshop . 2017

机译：边听边说：深度学习的语音链
5. Deep learning for speech classification and speaker recognition [D] . Saleem, Muhammad Muneeb. 2014

机译：深度学习用于语音分类和说话人识别
6. Speech Audiometry at Home: Automated Listening Tests via Smart Speakers With Normal-Hearing and Hearing-Impaired Listeners [O] . Jasper Ooster, Melanie Krueger, Jörg-Hendrik Bach, 2020

机译：主页言语听力测量：通过智能扬声器自动聆听测试具有正常听力和听力受损的听众
7. SPEAK YOUR MIND! Towards Imagined Speech Recognition with Hierarchical Deep Learning [O] . Pramit Saha, Muhammad Abdul-Mageed, Sidney Fels 2019

机译：说出你的想法！与分层深度学习的想象的语音识别

Listening while Speaking: Speech Chain by Deep Learning

摘要

著录项

相似文献

相关主题

期刊订阅